BJGP Open
● Royal College of General Practitioners
Preprints posted in the last 7 days, ranked by how well they match BJGP Open's content profile, based on 13 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Marban-Castro, E.; Muhwava, L.; Girdwood, S.; Kemp, T.; Freitas, J.; Kamau, Y.; Otieno, M.; Akach, D.; Morato, A.; Sanz, S.; Fiechter, V.; Erkosar, B.; Watson, M.; Vetter, B.; Haldane, C.; Shilton, S.; Rheeder, P.; Dave, J. A.; Carrihill, M.; Karsas, M.
Show abstract
Introduction: Continuous glucose monitoring (CGM) offers an advancement over traditional self-monitoring of blood glucose (SMBG) for people living with type 1 diabetes (T1D). However, evidence on the acceptability and feasibility of different CGM use cases in African populations remains limited. Methods: This was a pragmatic three-arm, randomised controlled trial on CGM conducted among people living with T1D in three public healthcare clinics in South Africa. Participants were assigned to Arm 1 (continuous CGM), Arm 2 (periodic CGM), or Arm 3 (SMBG). Diabetes education was provided at all study visits. Feasibility was assessed by adherence to CGM use and through the Glucose Monitoring Satisfaction Survey (GMSS). Diabetes distress was measured by the Diabetes Distress Scale (DDS), health-related quality of life (HRQoL) by the EQ-5D scales, and acceptability using the Theoretical Framework of Acceptability (TFA). Surveys were collected on paper and transferred to OpenClinica. Analyses were performed in R. The trial was registered in the Clinical Trials Registry (NCT05944718) on July 13, 2023. Results: A total of 83 participants were included in Arm 1, 85 in Arm 2, and 80 in Arm 3. CGM mean active time was 55% in Arm 1 versus 69% in Arm 2. The proportion of participants meeting the [≥]70% active time threshold was higher in Arm 2 (52%) than in Arm 1 (34%). Diabetes' distress declined across arms during the intervention period, with no significant difference between arms; distress increased slightly six months post-intervention but remained below baseline. At 6 months, glucose monitoring satisfaction was significantly higher in both CGM arms than in the SMBG arm, and satisfaction increased over time in CGM arms. Health-related quality of life remained stable across arms during the intervention period with no significant difference between arms. High acceptability was observed in both CGM arms, with higher ratings in the periodic arm. Conclusions: CGM was acceptable to people living with type 1 diabetes and feasible to use in public-sector clinics in South Africa, with high acceptability under continuous and periodic use. Health-related quality of life remained stable across arms, and diabetes-related distress declined, during the intervention period, across arms. Glucose monitoring satisfaction rose significantly in both CGM arms compared to SMBG. Periodic CGM might be a promising and potentially more scalable option than continuous use for public-sector care.
Chatzilena, A.; Hyams, C.; Challen, R.; Lahuerta, M.; McGuinness, S.; Clout, M.; Begier, E.; King, J.; Morales-Aza, B.; Duale, K.; Rodriguez Pereira, A.; Healy, W.; Southern, J.; Wells, P.; Lihou, K.; Grimes, C.; Campling, J. A.; Maskell, N.; Oliver, J.; Vyse, A.; Gessner, B.; Finn, A.; Danon, L.; The AvonCAP Research Group,
Show abstract
Introduction Acute lower respiratory tract disease (aLRTD) is a leading cause of hospitalisation and death, particularly in older adults and adults with comorbidities, with acute lower respiratory tract infection (aLRTI; pneumonia and non-pneumonic LRTI) being a major component. Non-pulmonary complications and functional decline after aLRTI are recognised, but their pathogen-specific burden is poorly described. We aimed to quantify renal, hepatic, thromboembolic and functional complications, and mortality, after aLRTI hospitalisation, by clinical phenotype and pathogen. Methods We conducted a cohort study of adults (>18 years) admitted with aLRTD to two hospitals in Bristol, UK (01 August 2022-31 July 2024). aLRTD was classified as pneumonia, non-pneumonic LRTI (NP-LRTI) or no diagnosis of aLRTI. Pathogens were identified from standard-of-care and research microbiology. Outcomes were acute kidney injury (AKI), acute liver dysfunction, venous thromboembolism (VTE), in-hospital falls, reduced mobility at discharge, increased care requirements, and 30-day and 1-year mortality. Analyses were descriptive. Results Among 246,797 adult admissions, 21,456 aLRTD hospitalisations were included: 10,239 (47.7%) pneumonia, 7,742 (36.1%) NP-LRTI and 3,475 (16.2%) with no evidence of aLRTI. Of 19,152 tested aLRTD admissions, 8,503 (44.4%) had a positive microbiological/virological test, yielding 9,204 pathogen detections; 1,194 (6.2%) had co-infections, and SARS-CoV-2 was most frequent, with influenza the second most common in pneumonia and NP-LRTI. Pneumonia had greater severity than NP-LRTI and no diagnosis of aLRTI (median length of stay 6 vs 4 vs 4 days; ICU admission 3.4% vs 0.7% vs 0.5%, respectively). Overall, 22.2% developed AKI, 6.1% acute liver dysfunction, 0.6% DVT and 2.4% PE; 1.8% had a fall, 11.5% reduced mobility, and 16.6% required increased care at discharge. 30-day and 1-year mortality were highest for pneumonia (14.0% and 32.0%, respectively). Pathogen-specific analyses showed longer stays and higher complications and mortality rates for SARS-CoV-2 and Streptococcus pneumoniae, and shorter stays with lower complication and mortality rates for influenza and Haemophilus influenzae. Conclusions Non-cardiovascular complications and functional decline after aLRTI were common, particularly in pneumonic and SARS-CoV-2 or pneumococcal disease. These findings support routine surveillance for renal, hepatic, thromboembolic events, early mobilisation and rehabilitation, and consideration of multi-system outcomes when evaluating public health and economic value of vaccines and therapies.
Amolo, P.; Mungai, L.; Karume, A. K.; Kibugi, J.; Mwende, W.; Botella, N.; Haldane, C.; Kamau, Y.; Marban-Castro, E.
Show abstract
Introduction Continuous Glucose Monitoring (CGM) is considered standard care in high-income countries. There is, however, limited published evidence on CGM use in low- and middle-income countries. The purpose of this study was to assess the usability, acceptability, and feasibility of CGM use among people living with type 1 diabetes (T1D) and caregivers in a low-resource setting. Research Design and Methods This prospective study conducted at the Kenyatta National Hospital purposively enrolled persons aged 4-25 years who had been on management for T1D for at least six months, and caregivers of those under 18 years. Fourty youth living with T1D used CGM for three months in place of self monitoring of blood glucose (SMBG). The System Usability Scale (SUS), a Theoretical Framework of Acceptability-based questionnaire, the Diabetes Distress Scale (DDS), the Glucose Monitoring Satisfaction Survey (GMSS), and a feasibility survey were administered. Outcomes were summarized descriptively, including means, medians, and frequencies using R statistical software. Results The median SUS score was 98.8 (IQR 92.5-100.0). Acceptability was high, and the median total GMSS score improved from 3.73 to 4.73. Among adolescents and adults, the median overall DDS score reduced from 1.54 to 1.36, with reductions in scores in all domains, except for hypoglycemia distress which increased, and physician distress which remained low. Among caregivers, the median overall DDS score declined from 2.05 (moderate distress) to 1.90 (low distress), with modest reductions in teen management and parent-teen relationship distress and a slight increase in personal distress. Median CGM active wear time was 89%. Conclusion This study comprehensively evaluated CGM across usability, acceptability, and feasibility outcomes, with the findings supporting the integration of CGM into routine diabetes management in low-resource settings. The short follow-up period, however, may not capture changing perceptions or long-term adherence.
Jaber, A.; Hughes, L.; Cameron, A. C.; Quinn, T. J.
Show abstract
Background: Systematic reviews of clinical prediction models increasingly include studies using artificial intelligence (AI) and machine learning (ML) methods alongside traditional multivariable regression approaches. A previously published Excel tool enabled standardised data extraction using the CHARMS checklist and risk of bias assessment using PROBAST. The recent publication of the PROBAST+AI framework, which distinguishes the assessment of model development quality from the assessment of model evaluation risk of bias and assesses applicability in both parts, necessitates an updated digital instrument applicable across prediction modelling methods. Methods: We updated an open-access Excel tool to incorporate the full PROBAST+AI framework. The updated template incorporates structural separation between assessment of model development quality and model evaluation risk of bias, with applicability assessed in both parts. It also incorporates updated signalling questions, including those addressing methodological issues particularly relevant to AI/ML, and automates the generation of summary tables and graphical displays. Results: The updated tool (CHARMS & PROBAST+AI Template) contains 11 worksheets and supports data extraction and appraisal for up to 30 prediction models. Dedicated, linked worksheets enable separate assessment of model development and model evaluation, with Domain 4 distinguishing among Apparent, Internal, and External evaluation settings. Key updates include dedicated assessments for predictor pre-processing, class imbalance handling and recalibration, data leakage prevention, and replication of the full model development pipeline within resampling procedures. Automated sheets dynamically format tables and summary charts covering PROBAST+AI parts. Conclusions: The CHARMS & PROBAST+AI Excel template provides a standardised, user-friendly, and rigorous digital framework for systematic reviewers appraising traditional statistical and AI-driven clinical prediction models.
Pinedo-Torres, I.; Taype-Rondan, A.; Vera-Luza, A. A.; Zegarra-Lizana, P. A.; Rojas-Vilca, J. L.; Yovera-Aldana, M.
Show abstract
Objective. To determine the publication rate of abstracts presented at the American Diabetes Association Scientific Sessions and to evaluate the association between statistical significance of study results and subsequent publication. Research Design and Methods. We conducted a retrospective cohort study of abstracts presented at the 2018 American Diabetes Association Scientific Sessions. The primary exposure was study result category (statistically significant vs. non-statistically significant findings), and the primary outcome was publication in an indexed journal within 5 years after conference presentation. Publication status was determined through PubMed/MEDLINE and Scopus searches. Adjusted relative risks (RRs) and 95% CIs were estimated using generalized linear models with Poisson distribution and robust variance. Results. Among 541 included abstracts, 321 (59.3%) were subsequently published in indexed journals. Abstracts reporting statistically significant findings had a higher publication rate than those reporting non-statistically significant findings (61.9% vs. 42.3%; p=0.002). In the adjusted analysis, abstracts with non-statistically significant findings had a lower likelihood of publication compared with those reporting statistically significant findings (adjusted RR 0.71 [95% CI 0.55-0.93]; p=0.013). Conclusions. Approximately four in ten abstracts presented at the ADA Scientific Sessions were not published within 5 years. Abstracts reporting non-statistically significant findings had a lower likelihood of subsequent publication, suggesting persistent publication bias in diabetology research. Future initiatives promoting the interpretation of effect estimates, confidence intervals and clinical relevance, rather than statistical significance alone, may help reduce selective dissemination of evidence
Sierpe, A.; Yen, R. W.; Milliman, A.; Cady, E.; Ahn, B.; Dade, A. E.; Devito, A. M.; Eckert, B. A.; Gopalan, V. V.; Krasinski, S. C.; MacMartin, M. A.; Musacchio, S. G.; Zhang, J.; Saunders, C. H.
Show abstract
Background Agenda-setting is a fundamental patient-centered communication practice in which a clinician works with a patient to elicit, propose, and organize topics for discussion during a clinical encounter. Various agenda-setting interventions have been developed, including patient-facing tools and clinician training, but their effects have not been systematically evaluated. We aimed to determine the effects of these interventions on encounter, patient, care partner, and clinician outcomes. Methods We searched grey literature and seven databases, including PubMed, from inception through July 2025 for randomized and non-randomized comparative studies of interventions designed to promote or improve clinical visit agenda-setting. Two reviewers independently screened articles and extracted data, with a third reviewer resolving conflicts. We assessed risk of bias using RoB 2 for randomized studies and ROBINS-I for non-randomized studies. We conducted random effects meta-analyses when outcomes were sufficiently comparable, assessed heterogeneity using I2, and rated certainty of evidence using GRADE. Post hoc exploratory subgroup analyses examined study design, adjustment status, and intervention structure. Results Twenty-nine articles describing 22 unique studies met the inclusion criteria, including 13 randomized and nine non-randomized studies. Agenda-setting interventions increased the occurrence of agenda-setting (risk ratio 5.43, 95% confidence interval (CI) 2.06 to 14.28, I2=34.6%) and favored the intervention for concerns addressed when measured as a continuous outcome (standardized mean difference (SMD) 0.37, 95% CI 0.16 to 0.57, I2=65.3%) and overall clinician satisfaction (SMD 0.50, 95% CI 0.23 to 0.78, I2=0.0%). There were no clear differences in the number of concerns raised (mean difference (MD) 0.21, 95% CI -0.19 to 0.61, I2=59.6%), visit duration (MD 0.64 minutes, 95% CI -0.83 to 2.12, I2=51.4%), or overall patient satisfaction (SMD 0.05, 95% CI -0.05 to 0.15, I2=47.0%). Potentially important heterogeneity was present for four of these six outcomes. Post hoc exploratory subgroup analyses did not provide clear evidence that effects varied by study design, adjustment status, or intervention structure. Risk of bias was often high, serious, or critical, and certainty of evidence was low or very low for all pooled outcomes. Conclusions To our knowledge, this is the first comprehensive synthesis of clinical visit agenda-setting interventions. Such interventions may increase the occurrence of agenda-setting and the extent to which patient concerns are addressed without increasing visit length. However, the certainty of evidence was low or very low, and the available evidence does not establish a superior intervention structure.
Withanage, N. D.; Perera, S.; Athiththan, L.
Show abstract
Background: Lumbar disc herniation, with or without concomitant disc degeneration, is a major cause of lumbar radiculopathy and low back pain, which also a key public musculoskeletal disorder without an exact pathophysiology. Studies have suggested that inflammatory cells and biochemical markers of inflammation also play an important role in lumbar radiculopathy in addition to nerve compression. The aim of the present study was to assess the association of selected circulatory inflammatory markers (CRP, hs-CRP and E-selectin) in patients with lumbar disc herniation without radiological degeneration (LDH) and lumbar disc herniation with radiological degeneration (LDHD). Materials & methods: This case-control study included 208 participants, comprising 104 patients with lumbar disc pathology and 104 controls. Patients were further stratified into LDH (n=67) and LDHD (n=37). Serum CRP, hs-CRP and E-selectin concentrations were measured. Results: Among the patients, 35.6 % presented with LDHD while 64.4 % had only LDH. Significantly increased median hs-CRP (p<0.001) and CRP (p<0.001) were observed in patients groups compared to controls, while CRP showing a consistent independent association across the combined disease (OR=1.68, 95% CI=1.33-2.14, p<0.001), LDHD (OR=1.62, 95% CI=1.16-2.20, p=0.005) and LDH (OR=1.69, 95% CI=1.30-2.20, p<0.001) multivariable models. No significant difference was observed in serum E-selectin between the study groups. Multivariable models incorporating inflammatory and clinical variables demonstrated substantially greater discriminatory performance than individual biomarkers alone. Conclusion: Elevated circulating CRP and hs-CRP concentrations were associated with lumbar disc pathology, with CRP showing a consistent independent association across the combined disease, LDH and LDHD multivariable models, whereas E-selectin showed no significant association. Multivariable models incorporating inflammatory and clinical variables demonstrated greater discriminatory performance than individual biomarkers. These findings support a potential systemic inflammatory component in lumbar disc pathology, although the cross-sectional nature of the measurements does not establish causality or a local inflammatory response within the disc.
Chowdhury, A. R.; Chowdhury, B.
Show abstract
Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.
Kremer, P.; Schlicker, N.; Hasnaj, R.; Bamberger, J.; Witte, T.; Haase, I.; Mayr, A.; Schmidt, C.; Osteras, N.; Baraliakos, X.; Kuhn, S.; Krusche, M.; Knitza, J.
Show abstract
Objectives To evaluate whether access to a certified large language model (LLM)-based clinical decision support system improves physician diagnostic performance in rheumatology compared with conventional diagnostic resources alone. Methods In this multicentre, open-label, randomised controlled trial, 82 physicians from seven hospitals in two countries were randomised 1:1 to conventional diagnostic resources plus Prof. Valmed or conventional resources alone. Participants assessed three rheumatology vignettes before and after assistance. The primary outcome was top-1 diagnostic accuracy. Secondary outcomes included top-3 accuracy, diagnostic reasoning, confidence, case-processing time and perceived support quality. Results Top-1 accuracy increased from 22.2% to 33.3% in the intervention group and from 23.3% to 35.0% in the control group, with no between-group difference in improvement (adjusted OR 0.99, 95% CI 0.45 to 2.19; p=0.979). Differences in top-3 accuracy, diagnostic reasoning and confidence were also not significant. Assisted case-processing time was substantially shorter with LLM support (94 vs 206 s; adjusted mean difference -112 s, 95% CI -141 to -83; p<0.001). Information timeliness and perceived diagnostic support quality were rated significantly higher in the intervention group. Exploratory analyses showed persistent overconfidence and substantial AI over-reliance. Conclusions Certified LLM-based diagnostic support did not improve diagnostic accuracy compared with conventional resources, but substantially reduced case-processing time and improved perceived support quality. These findings suggest potential workflow benefits while highlighting overconfidence and over-reliance as important safety considerations.
Chaturvedi, R. R.; Gracner, T.; Perez-Arce, F.; Suen, S.-c.; Jin, J.; Orriens, B.; Pacula, R. L.; Sexton Ward, A.; Haile, R.; Kapteyn, A.
Show abstract
Importance: Evidence on GLP-1/GIP therapies is largely derived from trials enrolling selected populations or medical records that miss utilization outside healthcare channels. No nationally representative cohort has characterized real-world uptake, indications, and access. Objective: To characterize GLP-1/GIP prevalence, indication, clinical profile, and access. Design: Prospective cohort study with three GLP-1/GIP surveillance waves (March 2024, December 2024, October 2025). Setting: The Understanding America Study, an address-based, nationally representative panel of approximately 15,000 US adults aged 18+ years initiated in 2014. Participants: UAS participants responding to at least one surveillance wave (n=9150). Exposures: GLP-1/GIP use status (never vs any use, comprising current and former use), self-reported primary indication (diabetes, weight loss, or other), and access pathway (traditional vs non-traditional). Main Outcomes and Measures: Survey-weighted prevalence of GLP-1/GIP use, overall and by indication and access pathway; sociodemographic, cardiometabolic, treatment, and access characteristics; and smartwatch-derived resting heart rate, heart rate variability, maximum activity heart rate, step count, and sleep duration and variability. Results: Among n=9150 adults (1274 with any use; 60.9% female; median age 53 years), weighted prevalence increased 46%, from 8.2% (March 2024) to 12.0% (October 2025) representing 32 million. Weight-loss indications grew, reaching nearly half of use (4.1% to 5.6%); diabetes-indicated use was stable (5.3% to 5.4%). Users carried high cardiometabolic burden (obesity, 68.2%; diabetes, 53.6%) but diverged by indication: diabetes-indicated users were older (median, 59 vs 49 years), whereas weight-loss-indicated users were more often female (69.9% vs 51.3%) and healthier. One in three users (~9 million) had non-traditional access, especially in weight-loss-indicated users, of whom 33% had no conventional prescription; 41% used compounding, online, or foreign pharmacies; and, 43% lacked coverage. Non-traditional users were five times as likely to report an unlisted, likely compounded formulation (19.8% vs 4.1%). All p<0.05. Conclusions and Relevance: Real-world GLP-1/GIP use has grown rapidly and diversified substantially in indication, access, and population profile. One in 3 users obtained treatment through nontraditional channels largely invisible to claims data, raising long-term safety, efficacy, and coverage questions. GLIMMER provides a public, nationally representative longitudinal evidence base for future payer and provider decisions.
Khan, Z.; McCarthy, C.; Dalton, K.; Jungo, K. T.; Doherty, A. S.; Reeve, E.; Moriarty, F.
Show abstract
Background: Adverse drug withdrawal events (ADWEs) are a key safety concern during deprescribing but remain poorly explored in pharmacovigilance systems. Objectives: To identify and compare ADWE signals across drug classes, different drugs within drug classes, and across patient characteristics, countries, and over time. Methods: A case/non-case disproportionality analysis was conducted in FDA-FAERS and EMA-EudraVigilance pharmacovigilance databases, with stratification by age (adults: 18-64, older adults: [≥]65), sex (male/female), reporting time (2004-2023 in 5-year intervals), and country (for EMA data). Disproportionality analysis (quantitative signal detection) was used to detect signals between ADWEs and drugs using the proportional reporting rate (PRR[≥]2), reporting odds ratio (ROR>1), and information component (IC>0) with case count [≥]5. Results: Overall, 158,501 reports (FDA-FAERS 145,514; EMA-EudraVigilance 12,987) included drug-event pairs related to ADWEs. In FDA-FAERS, clobetasone (IC=5.58; PRR=79.18; ROR=176.90) showed the strongest ADWE signals, followed by hydromorphone (4.85; 29.94; 37.37), hydrocodone, and paroxetine. In EMA-EudraVigilance, ethyl loflazepate (IC=6.01; PRR=119.80; ROR=197.53), clobetasone (5.39; 102.73; 155.10), veralipride, and levomethadone had the strongest signals. Most drugs maintained positive ADWE signals in analysis stratified into adults and older adults. However, among the top 10 drugs (based on highest IC values), buprenorphine/naloxone, desvenlafaxine, and baclofen in FDA-FAERS (ICs 4.95-6.05) showed stronger signals in older adults. A sex-based difference was observed, with paroxetine, venlafaxine, and buprenorphine/naloxone showing a stronger positive signal in females in both databases, whereas several opioids had stronger signals in males versus females across both databases. Conclusion: This study suggests ADWE signals for some medications differ by age and sex, potentially indicating different risks for withdrawal effects.
Greendyk, J. D.; Allen, W. E.; Hossain, A.; Trichas, Z.
Show abstract
Background: Percutaneous mechanical circulatory support (pMCS) is increasingly used in critically ill patients, yet its value in relation to cost and outcomes remains unclear. We evaluated national variation in utilization, outcomes, and cost, and introduced a value of care framework integrating risk-adjusted outcomes and expenditures. Methods: We performed a retrospective cohort study using the National Inpatient Sample to identify non-elective hospitalizations of critically ill patients undergoing intra-aortic balloon pump (IABP) or percutaneous left ventricular assist device (pLVAD) placement using ICD-10 codes. Multivariable logistic regression and generalized linear models were used to estimate expected outcomes and costs. Observed-to-expected (O/E) ratios were calculated, and a value index was derived to compare procedural strategies. Results: A total of 57,910 weighted hospitalizations were included (IABP 78%, pLVAD 22%). In-hospital mortality exceeded 30% across regions. Significant regional variation was observed, with the West demonstrating the highest costs and the Midwest the lowest (p<0.001). Mean hospital charges were higher for pLVAD compared with IABP ($403,731 vs $320,769). Both strategies achieved outcomes better than expected after risk adjustment (O/E 0.92); however, costs were higher than expected for both, with greater relative cost inflation observed for IABP (O/E 1.41) and higher absolute costs for pLVAD. In value-of-care analysis, IABP was associated with lower cost and comparable outcomes, while pLVAD demonstrated higher cost without proportional outcome improvement. Conclusion: Substantial variation exists in the cost, outcomes, and value of pMCS strategies. While both IABP and pLVAD achieve favorable risk-adjusted outcomes, pLVAD is associated with higher costs without commensurate clinical benefit.
Brodtmann, A.; Patel, S.; Restrepo, C.; Khlif, M. S.; Werden, E.; Ellis, R.; Alsawaf, S.; Ekinci, E. I.; Srivastava, P. M.; Ramchand, J.; MacIsaac, R. J.; Churilov, L.; Burrell, L. M.
Show abstract
BACKGROUND People with type 2 diabetes mellitus (T2DM) are at higher risk of cerebral small vessel disease and left ventricular hypertrophy (LVH), potentially contributing to cognitive decline and dementia. We aimed to describe brain volume and cognitive trajectories over 2 years in a cohort of people with T2DM and to determine whether LVH causes increased brain atrophy and cognitive decline. METHODS Diabetes and Dementia (D2) study is a multicentre observational cohort study in Melbourne, Australia. Participants aged >50 years were recruited via 2 hospital outpatient clinics, 3 private clinics, and study advertisements. Participants with pre-existing cognitive impairment, life-limiting medical illness, and severe chronic renal impairment were excluded. Participants attended study visits for brain MRI, transthoracic echocardiography (TTE), and cognitive testing at baseline and 2 years. The exposure was LVH determined on baseline TTE. Pre-specified outcomes were total brain volume (TBV) change and cognitive decline (z-score change?-1 in any cognitive domain) over 2 years. Regression analyses examined associations between baseline variables and outcomes. A causal inference approach was utilized using inverse probability of treatment weighting to standardize for confounding covariates, excluding participants for non-positivity on age and baseline TBV. RESULTS Participants were recruited 20May2016 to 20March2020: 2378 screened, 702 eligible, 196 consented, 150 baseline and 123 2-year assessments with complete MRI, TTE, and cognitive data (17.4% attrition). At baseline, LVH was associated with female sex, older age, lower educational attainment, lower mood, hypertension, obesity, beta-blocker use, and smaller TBV. Participants with baseline cognitive impairment exhibited greater brain atrophy. Lower educational attainment, hypertension, and lower baseline cognitive scores were associated with cognitive decline. Causal inference analysis included 62 participants with no LVH (20(32%) women; mean [SD]=66.9[5.9] years), and 31 with LVH (17(55%) women, 67.4[5.4] years). LVH caused lower TBV change: standardized mean difference (95% CI) 6.3 (0.1, 12.5) cm3, P=.048. LVH had no effect on cognitive decline. CONCLUSIONS Brain atrophy and cognitive decline were associated with baseline cognitive impairment. LVH caused less brain atrophy and cognitive decline in people with T2DM. We conclude that guideline-directed LVH therapies such as beta-blockers have both cardioprotective (remodelling) and neuroprotective effects. TRIAL REGISTRATION ACTRN12616000546459 UTN: U1111-1181-6659
Witham, M.; Evison, F.; Bellass, S.; Cooper, R.; Gallier, S.; Pretorius, S.; Sapey, E.; Suklan, J.; Sayer, A. A.
Show abstract
Study Objective Little is known about where in hospital care for multiple long-term conditions (MLTC) is delivered. We aimed to describe pathways of care (ward transfers) and outcomes for people admitted to hospital for unscheduled care by MLTC status and other key sociodemographic characteristics. Design and setting Analysis of routinely-collected electronic health records from a large acute UK hospital. Participants Adult unscheduled care admissions from 1st July 2018 to 30th June 2019. The presence of two or more of 59 long-term conditions was ascertained using ICD-10 codes from previous hospital discharges. Main outcome measures Markov state transition probabilities were derived for ward moves and compared for MLTC vs no MLTC, age, sex, ethnicity and neighbourhood deprivation. Outcomes (length of stay, death, readmission, move from definitive ward) and time spent in emergency and assessment departments were compared between subgroups. Results A total of 33,252 adults, mean age 56.0 (SD 21.9) years were analysed; 14,834 (42.4%) had MLTC. People with MLTC were more likely to die in hospital (4.2 vs 1.9%, p<0.001), transfer to internal medicine wards or older peoples medicine wards, were less likely to transfer to surgical wards, had longer median length of stay (1.83 vs 0.69 days, p<0.001), stayed longer in acute medical units (15.5 vs 9.6 hours, p<0.001), and were more likely to move from their definitive ward (18.2 vs 16.4%, p=0.002). Conclusion Unscheduled hospital care pathways are complex and differ for people with MLTC, who have worse outcomes and may be less likely to receive optimal care.
Dyer, B. P.; Deery, M.; Heyman, R.; Robinson, P.; Wainwright, C.; Sly, P.; Ware, R.; Blake, T.
Show abstract
Background Elexacaftor-tezacaftor-ivacaftor (ETI) has been demonstrated to improve lung function in clinical trials; however, evidence describing effects on trajectories and whether long-term improvements are sustained (>1-year) is lacking. We estimated within-person lung clearance index (LCI) trajectories before and after ETI initiation, assessing changes in level and rate of change, alongside acute LCI change, up to three years after ETI initiation. Methods Prospective observational study of children at a tertiary hospital. Children aged 3-17 years with [≥]2 LCI testing occasions (i) before and (ii) after starting ETI were used to describe lung function trajectories. Children with [≥]1 pre-ETI and [≥]1 post-ETI LCI occasion(s) were used to describe acute LCI change after ETI initiation. Age-adjusted LCI trajectories for time periods (i) before and (ii) after ETI initiation were estimated using linear mixed-effects models, and pre- and post-ETI LCIs were compared using paired Wilcoxon tests. Results Mean pre-ETI and post-ETI longitudinal changes in LCI were -0.007 (95% CI: -0.28, 0.27; n=35) and 0.12 (95% CI: -0.17, 0.41; n=20) turnovers per year, respectively. Before ETI initiation, 57% (30/53) of patients had an LCI[≥]7.1 turnovers (indicating impaired lung function), compared to 26% (14/53) post-ETI, with a median LCI difference of -0.70 (95% CI -0.84, -0.46; p<0.001) turnovers. Within-individual variability in LCI decreased post-ETI. Conclusions Our real-world data within a unique longitudinal study provide a comprehensive picture of ETI benefit by outlining not only acute improvement in LCI but maintained stability in LCI trajectories and improved LCI stability sustained up to three years post-initiation.
Oyarzun-Silva, R. A.; Hernandez-Hernandez, P.; Fernandez-Vaquero, M. A.; De Luis-Cabezon, N.
Show abstract
Background. Videolaryngoscopy still requires adjuncts or hyperangulated rescue in a clinically important minority, and bedside screening discriminates modestly. Point-of-care ultrasound (POCUS) of the anterior airway is a promising alternative, but existing prediction models are opaque or assume a pre-specified functional form. We developed and internally validated a parsimonious, fully disclosed POCUS risk equation whose form is recovered from data and whose structural properties are machine-checked by formal proof - to our knowledge the first formally verified clinical risk predictor - following TRIPOD+AI 2024. Methods. In a prospective single-centre, single-operator cohort of 259 adults undergoing elective videolaryngoscopy (no-Easy airway 68/259, 26.3%), Sequentially Thresholded Least Squares with bootstrap stability selection (B=300) screened a 71-term library of nine POCUS features and retained a seven-term logistic equation; a two-term bootstrap-stable model was pre-specified as robustness analysis. Internal validation used 5x10 repeated cross-validation plus temporal and device hold-outs, with pre-specified overfitting and optimism assessments. Five behavioural properties of the deployed equation were machine-checked in Lean 4. Results. Two interactions met the |c|/sigma_c>2 stability criterion: skin-to-epiglottis x skin-to-hyoid-bone distance and tongue volume x sagittal tongue area. The seven-term equation reached a 5x10 cross-validated C-statistic of 0.966 (optimism-corrected 0.968) and held across temporal and device hold-outs (0.94-0.97). Calibration-in-the-large matched prevalence, with cross-validated slope 0.90 attenuating to 0.625 out-of-time; standard recalibration restored 0.92 without loss of discrimination. The pre-specified two-term robustness model reproduced this performance (C-statistic 0.964-0.968; events-per-parameter 34; shrinkage 0.99), confirming the result is not an artefact of the screening stage. Net benefit over a clinical baseline was positive across 10-50% thresholds. All five Lean 4 theorems compiled without sorry. Conclusions. A sparse, formally verified POCUS equation predicts difficult videolaryngoscopy with high internally validated discrimination and quantified, modest overfitting. Because the equation was developed in a single-operator cohort and its inputs are operator-dependent, external validation requires prior harmonisation of the measurement protocol and operator credentialing.
Mathew, Z.; Mehta, R.; Kim, S.; Jeyaraj, J.; Asif, T.
Show abstract
Background: Primary malignant cardiac tumors (PMCTs) are rare and histologically heterogeneous. Objective: To compare demographics, specific ICD-O-3 morphologies, first-course treatment patterns, annual registered case counts, and unadjusted overall survival between soft-tissue and hematologic PMCTs. Methods: We identified 730 PMCT cases diagnosed from 2000 to 2021 in SEER 18 (ICD-O-3 topography C38.0). Histologic lineage was assigned from ICD-O-3 morphology. Comparative analyses included soft-tissue (n=458) and hematologic (n=212) tumors. First-course variables were primary-site surgery, chemotherapy (yes versus no/unknown), and radiotherapy (radiation versus none/unknown). Groups were compared with chi-square tests. Overall survival was estimated with Kaplan-Meier methods; follow-up was truncated at 120 months. Results: Soft-tissue PMCTs occurred predominantly at ages 45-64 years (67.9%), whereas hematologic PMCTs occurred predominantly at age [≥]65 years (63.2%; p<0.001). Men comprised 59.9% of hematologic and 49.3% of soft-tissue cases (p=0.014). The leading soft-tissue morphology was hemangiosarcoma/angiosarcoma (ICD-O-3 9120/3; 201/458, 43.9%); synovial sarcoma accounted for 20/458 cases (4.4%). Diffuse large B-cell lymphoma, NOS, accounted for 131/212 hematologic tumors (61.8%). Any primary-site surgery was recorded in 66.6% of soft-tissue versus 15.6% of hematologic cases (p<0.001). Chemotherapy was recorded in 67.5% versus 51.1% (p<0.001), and radiotherapy in 9.0% versus 20.5% (p<0.001). In exploratory Kaplan-Meier analyses, hematologic patients with recorded chemotherapy had higher unadjusted 120-month overall survival than those without recorded chemotherapy (42.0% versus 12.2%; log-rank p=7.5x10-). Radiation-associated survival differences were not statistically significant in either lineage. Conclusions: Soft-tissue and hematologic PMCTs have distinct age distributions, named histologies, and first-course treatment patterns in SEER. These findings describe registry coding and do not establish treatment effectiveness or population incidence.
Green, J. L.; Davies, H.; Russell, D. A.
Show abstract
Background: The relative merits of infrainguinal bypass and primary major lower limb amputation (MLLA) for chronic limb-threatening ischaemia (CLTI) remain uncertain, and the baseline profiles of patients selected for each strategy are poorly described. Methods: A systematic review and meta-analysis were undertaken in accordance with PRISMA 2020 and prospectively registered (PROSPERO: CRD42022356094). MEDLINE, Embase, CENTRAL, and CINAHL were searched from inception to March 2025. Prospective studies of adults with CLTI undergoing primary infrainguinal bypass or primary MLLA were eligible. Mortality, major adverse cardiovascular events (MACE) and subsequent amputation outcomes were synthesised using random-effects meta-analysis of proportions. Baseline comorbidity profiles were also extracted. Results: Twenty-seven studies involving 6,576 patients were included: 5,779 underwent infrainguinal bypass and 797 underwent MLLA. After bypass, pooled mortality was 3.7% at 30 days (95% CI 2.8%-4.9%, I2 = 49.4%), 18.5% at 1 year (95% CI 15.6%-21.9%, I2 = 62.3%), and 54.3% at 5 years (95% CI 50.5%-58.0%, I2 = 0%). After MLLA, pooled mortality was 9.2% at 30 days (95% CI 4.1%-19.3%, I2 = 73.5%), 28.5% at 1 year (95% CI 13.3%-51.0, I2 = 70.8%), and 39.9% at 2 years (95% CI 0.3%-99.3, I2 = 90.5%), although longer-term estimates were limited by sparse data and marked heterogeneity. Thirty-day MACE was 6.5% (95% CI 4.3%-9.7, I2 = 63.5%) after bypass and 2.8% after MLLA (95% CI 0.1%-37.6%, I2 = 0%). Early subsequent major amputation after bypass occurred in 3.9% of patients (95% CI 2.0%-7.7%, I2 = 91.2%), rising to 16.2% at 1 year (95% CI 12.6%-20.5%, I2 = 82.0%) and 33.3% at 3 years (95% CI 20.1%-49.8%, I2 = 0%). Early re-amputation after MLLA occurred in 10.9% of patients (95% CI 4.5%-24.4%, I2 = 40.3%). Baseline comorbidity burden was high in both groups, with substantial heterogeneity across studies. Conclusions: CLTI carries a poor prognosis regardless of treatment strategy. Infrainguinal bypass is associated with lower early mortality and better early limb preservation than primary MLLA, but long-term survival remains poor and later limb failure is common. Primary MLLA is not a low-risk alternative. Better contemporary comparative evidence utilising modern causal inference approaches is needed to support individualised decision-making.
Li, Z.; Fujisawa, T.; Skadberg, O.; Fineran, P.; Thurston, A. J.; Tew, Y. Y.; Aakre, K. M.; Mills, N. L.; Wereski, R.; the POC-ET Investigators,
Show abstract
Background: High-sensitivity cardiac troponin (hs-cTn) assays enable safe early discharge of patients at very low risk for myocardial infarction. We previously developed a single-sample rule-out pathway using the ARCHITECT hs-cTnI assay to risk stratify patients with suspected acute coronary syndrome. In a secondary analysis of the POC-ET (Point of Care Evaluation of High-sensitivity Cardiac Troponin) study, we evaluated performance of risk stratification with the Alinity hs-cTnI assay. Methods: Patients presenting with possible myocardial infarction in the POC-ET (NCT05665127) study were included. The primary outcome was type 1, 4b or 4c myocardial infarction or cardiac death at 30 days. Cardiac troponin I (cTnI) was measured in stored materials using the ARCHITECT and Alinity hs-cTnI assays. The sex-specific 99th percentile upper reference limit (URL) are 34 ng/L in men and 16 ng/L in women for both assays. Agreement was assessed with Bland-and-Altman limit of agreement method, Passing Bablok regression, and Pearson's correlation coefficient. Distributions of presentation measurements were compared with Kolmogorov-Smirnov test. Performance was evaluated in the overall population and prespecified subgroups. The negative predictive value (NPV) and sensitivity were determined and proportion of patients identified as low, intermediate, and high risk were calculated and modelled using ordinal logistic regression. Results: In 986 patients (60 [51-70] years, 38% female), 78 (7.9%) had a primary outcome. Strong agreement was found in the raw cTnI measurements (99% samples within the Bland-Altman limit of agreement; correlation coefficient r: 0.967 (95% CI 0.964-0.969, P<0.001); Passing Bablok regression: slope 1.12 [1.11-1.13], intercept -0.16 [-0.18 to -0.13]). At presentation, distributions of cTnI measurements by the two assays were similar (P=0.810). Both assays showed comparable diagnostic performance using a risk stratification threshold of <5 ng/L and the sex-specific diagnostic threshold, with the same NPV (Alinity 100 [99.7-100]% versus ARCHITECT 100 [99.7-100]%) and sensitivity (Alinity 100 [97.3-100]% versus ARCHITECT 100 [97.3-100]%). Similar proportions of patients stratified as low- (Alinity 67% versus ARCHITECT 67%), intermediate-risk (23% versus 24%) and high-risk (10% versus 9%) at presentation with minor reclassification. Similar efficacy was observed across subgroups stratified by sex, age, history of myocardial infarction, renal function, and symptom duration. Conclusions: The Alinity hs-cTnI and the ARCHITECT hs-cTnI assays can be used interchangeably in the assessment of suspected myocardial infarction with comparable safety and efficacy.
Dick, M.; Madathil, S.; Patel, A.; Kapoor, H. S.; Sharma, M.; D'Souza, Z.; Hameed, S.; Abu-Samak, M.; Najirad, A.; Dwairi, D.; Radaideh, O.; Nicolau, B.
Show abstract
Objectives: Dentists prescribe approximately one in ten antibiotics worldwide, yet antimicrobial stewardship (AMS) remains underemphasized in dental education. Large language models (LLMs) may support AMS training, but their proficiency and clinical reasoning in this context remain unclear. We evaluated GPT-4o's accuracy and clinical reasoning on dental antibiotic prescribing questions, stratified by question difficulty. Methods: We assembled 125 multiple-choice questions on dental antibiotic prescribing from eight peer-reviewed studies (2017-2023). GPT-4o answered each question and generated a clinical justification. Accuracy was assessed against source-study answer keys and examined across difficulty quartiles. Justifications were evaluated using an adapted 12-axis human-evaluation framework assessing scientific consensus, extent and likelihood of harm, inappropriate and missing content, bias, and both correct and incorrect comprehension, retrieval, and reasoning. Prophylaxis-specific questions were analysed separately. Results: GPT-4o correctly answered 72% of questions. Accuracy remained relatively stable across difficulty quartiles (78%, 78%, 65%, 70%). Experts rated 95.4% of justifications positively across the 12 axes. Comprehension, retrieval, and reasoning each exceeded 96.2% positive ratings. Missing content was the main weakness (7.8%), and 7.1% of justifications showed a moderate-to-severe potential for harm. Performance on prophylaxis-specific questions (98.1%) exceeded non-prophylaxis questions (93.0%). Conclusions: GPT-4o demonstrated moderate-to-high proficiency and clinically defensible reasoning in dental antibiotic prescribing questions. However, residual risks indicate that it is not suitable for unsupervised clinical use but shows potential as a supervised AMS educational tool.